Papers with AI models

36 papers
Geo-Cultural Representation and Inclusion in Language Technologies (2024.lrec-tutorials)

Copied to clipboard

Challenge: audi et al.: training and evaluation of language models rely on semi-structured data that is annotated by humans . e-learning tools do not integrate rich and diverse community perspectives into language technologies .
Approach: They will examine how different socio-cultural perspectives influence what is taken as ground truth by models.
Outcome: This tutorial examines how different socio-cultural perspectives influence representations of global concepts.
Human-AI Collaboration: How AIs Augment Human Teammates (2025.acl-tutorials)

Copied to clipboard

Challenge: Despite the potential of general-purpose models, they are far from perfect, excelling at certain tasks while struggling with others.
Approach: This tutorial will review recent developments related to human-AI teaming and collaboration.
Outcome: This tutorial will review recent developments related to human-AI teaming and collaboration.
Human-AI Interaction in the Age of LLMs (2024.naacl-tutorials)

Copied to clipboard

Challenge: Large Language Models (LLMs) have revolutionized the capabilities of AI systems.
Approach: This tutorial will provide an overview of the interaction between humans and Large Language Models (LLMs) it will start with a review of the types of AI models we interact with and walkthrough of the core concepts in Human-AI Interaction.
Outcome: This tutorial will provide an overview of the interaction between humans and LLMs, exploring the challenges, opportunities, and ethical considerations that arise in this dynamic landscape.
Transforming Brainwaves into Language: EEG Microstates Meet Text Embedding Models for Dementia Detection (2025.acl-srw)

Copied to clipboard

Challenge: Dementia is recognised as the seventh leading cause of mortality globally and plays a major role in increasing disability and dependence among older adults.
Approach: They propose to represent electroencephalography microstates as symbolic, language-like sequences and use text embedding and time-series deep learning models for classification.
Outcome: The proposed method achieves a high accuracy of 94.31% on 1001 EEG data from multiple countries and eliminates fixed configurations and costly/invasive modalities.
Do Androids Laugh at Electric Sheep? Humor “Understanding” Benchmarks from The New Yorker Caption Contest (2023.acl-long)

Copied to clipboard

Challenge: Large neural networks can generate jokes, but do they really “understand” humor? a new challenge challenges AI models to match a joke to a cartoon, identify a winning caption, and explain why a winner is funny.
Approach: They propose three tasks based on the New Yorker Cartoon Caption Contest . they aim to match a joke to a cartoon, identify a winning caption and explain why it's funny .
Outcome: The proposed tasks are based on the New Yorker Cartoon Caption Contest . they include matching a joke to a cartoon, identifying a winning caption, and explaining why a funny caption is funny.
Diverse Perspectives, Divergent Models: Cross-Cultural Evaluation of Depression Detection on Twitter (2024.naacl-short)

Copied to clipboard

Challenge: Social media data is used for detecting users with mental disorders, but public datasets lack crucial metadata related to this aspect.
Approach: They use a custom geo-located Twitter dataset to evaluate the generalization of depressiondetection models on cross-cultural Twitter data.
Outcome: The proposed models perform worse on Global South users compared to Global North.
VN-MTEB: Vietnamese Massive Text Embedding Benchmark (2026.findings-eacl)

Copied to clipboard

Challenge: a lack of large-scale test datasets makes it difficult to evaluate AI models before deploying them in real-world projects.
Approach: They propose a Vietnamese benchmark for embedding models that leverages large language models and embeddable models to translate and filter samples from the Massive Multilingual Text Embedding Benchmark.
Outcome: The proposed benchmark outperforms existing models in Vietnamese and English tasks with 41 datasets.
LLM-GEm: Large Language Model-Guided Prediction of People’s Empathy Levels towards Newspaper Article (2024.findings-eacl)

Copied to clipboard

Challenge: Empathy is a key component of human-to-human interactions, and is often overlooked due to the inherent noise in crowdsourced annotations.
Approach: They propose a large language model-guided empathy prediction system that rectifies annotation errors based on defined annotation selection threshold and makes annotations reliable for conventional empathy prediction models.
Outcome: The proposed system rectifies annotation errors based on defined selection threshold and makes the annotations reliable for conventional empathy prediction models, e.g., BERT-based pre-trained language models.
Generate then Select: Open-ended Visual Question Answering Guided by World Knowledge (2023.findings-acl)

Copied to clipboard

Challenge: Open-ended Visual Question Answering (VQA) requires models to reason over visual and natural language inputs using world knowledge.
Approach: They propose a new VQA pipeline that deploys a generate-then-select strategy guided by world knowledge for the first time.
Outcome: The proposed pipeline expands the knowledge coverage from in-domain training data by 4.1% on OK-VQA, without additional computation cost.
ARES: Alternating Reinforcement Learning and Supervised Fine-Tuning for Enhanced Multi-Modal Chain-of-Thought Reasoning Through Diverse AI Feedback (2024.emnlp-main)

Copied to clipboard

Challenge: Large Multimodal Models excel at comprehending human instructions and demonstrate remarkable results across a broad spectrum of tasks.
Approach: They propose an algorithm that alters REinforcement Learning and Supervised Fine-Tuning to refine large multimodal models with specific preferences.
Outcome: The proposed algorithm achieves 70% win rate compared to baseline models judged by GPT-4o.
Interactive Text Generation (2023.emnlp-main)

Copied to clipboard

Challenge: Advances in generative modeling have made it possible to automatically generate high-quality texts, code, and images, but they can be unsatisfactory in many respects.
Approach: They propose a task that allows training generation models interactively without the costs of involving real users.
Outcome: The proposed model trains with Imitation Learning without the cost of involving real users and is superior to non-interactive models.
SEACrowd: A Multilingual Multimodal Data Hub and Benchmark Suite for Southeast Asian Languages (2024.emnlp-main)

Copied to clipboard

Challenge: Southeast Asia (SEA) is home to over 1,300 indigenous languages and 671 million people . prevailing AI models suffer from a significant lack of representation of texts, images, and audio datasets from SEA .
Approach: They propose to provide a resource center that provides standardized corpora in nearly 1,000 SEA languages across three modalities.
Outcome: a new benchmark assesses the quality of AI models on 36 SEA languages across 13 tasks . the results highlight the importance of SEA as a culturally diverse region .
Adaptive Parameter Compression for Language Models (2025.findings-naacl)

Copied to clipboard

Challenge: Adaptive parameter compression is a new approach to improve NLP models . the current algorithm is based on a single parameter, but it is not scalable.
Approach: They propose a hardware-independent compression strategy that extends the weight-squeezing approach by introducing compression biases and weights.
Outcome: The proposed compression strategy outperforms DistilBERT base models while being significantly more efficient.
SenticNet 7: A Commonsense-based Neurosymbolic AI Framework for Explainable Sentiment Analysis (2022.lrec-1)

Copied to clipboard

Challenge: Despite recent advances, AI still struggles with complex tasks that require commonsense reasoning such as natural language understanding.
Approach: They propose a commonsense-based framework that aims to overcome these limitations in the context of sentiment analysis.
Outcome: The proposed framework overcomes these limitations in the context of sentiment analysis.
TOP-Training: Target-Oriented Pretraining for Medical Extractive Question Answering (2025.coling-main)

Copied to clipboard

Challenge: e-health records underscore the growing significance of information extraction (IE) from these datasets.
Approach: They propose a target-oriented pre-training paradigm for extractive question-answering in the medical domain . TOP-Training moves one step further than popular domain-oriented fine-tuning .
Outcome: The proposed method improves on the Medical-EQA benchmarks.
Efficient Unstructured Pruning of Mamba State-Space Models for Resource-Constrained Environments (2025.emnlp-main)

Copied to clipboard

Challenge: State-space models struggle with quadratic computational complexity, limiting their use in long-context tasks and resource-constrained input data.
Approach: They propose a pruning framework specifically tailored for Mamba that reduces parameter counts by 70% with only a 3–9% drop in performance.
Outcome: The proposed pruning framework achieves up to 70% parameter reduction with only a 3–9% drop in performance.
Where Fact Ends and Fairness Begins: Redefining AI Bias Evaluation through Cognitive Biases (2025.findings-emnlp)

Copied to clipboard

Challenge: Existing benchmarks conflate factual correctness and normative fairness . a model may generate responses that are factually accurate but socially unfair .
Approach: They propose a benchmark to examine the boundary between fact and fair . they draw on representativeness bias, attribution bias and ingroup–outgroup bias to explain why models often misalign fact and faireness.
Outcome: The proposed model is based on ten frontier models and is available on github . it is compared with a standard model that generates people of color in Nazi-era uniforms .
Bridging the Digital Divide: Performance Variation across Socio-Economic Factors in Vision-Language Models (2023.emnlp-main)

Copied to clipboard

Challenge: Among the minority groups under-represented in AI, data from low-income households are often overlooked in data collection and model evaluation.
Approach: They evaluate the performance of a vision-language model on a geo-diverse dataset . they highlight insights that can help mitigate these issues and propose actionable steps for economic-level inclusive AI development.
Outcome: The proposed model performs lower for the poorer groups than the wealthier groups across topics and countries.
StatsChartMWP: A Dataset for Evaluating Multimodal Mathematical Reasoning Abilities on Math Word Problems with Statistical Charts (2025.findings-emnlp)

Copied to clipboard

Challenge: StatsChartMWP is a dataset for evaluating visual mathematical reasoning abilities on math word problems with statistical charts.
Approach: They propose a dataset for evaluating visual mathematical reasoning abilities on math word problems with statistical charts.
Outcome: The proposed model is more effective than open-source approaches.
What is More Likely to Happen Next? Video-and-Language Future Event Prediction (2020.emnlp-main)

Copied to clipboard

Challenge: Existing models cannot make multimodal commonsense predictions of future events based on video and dialogue .
Approach: They propose a task to predict which event is more likely to happen in a video clip . they use a dataset with 28,726 future event prediction examples from 10,234 videos .
Outcome: The proposed model provides a good starting point but leaves room for future work.
ACQUIRED: A Dataset for Answering Counterfactual Questions In Real-Life Videos (2023.emnlp-main)

Copied to clipboard

Challenge: despite its importance, there are few datasets that cover multimodal counterfactual reasoning . a dataset focusing on this area is limited because of its limited coverage over synthetic environments .
Approach: They develop a video question answering dataset that provides questions on multimodal reasoning . they ask questions about counterfactual hypotheses over visual events .
Outcome: The proposed dataset shows a significant performance gap between models and humans . it provides questions that span physical, social, and temporal dimensions .
Evolving Agents (2026.acl-long)

Copied to clipboard

Challenge: Current models are static entities incapable of compressing complexity of real world into generalisable concepts . authors: lack of endogenous mechanism for representation updating renders models vulnerable to domain mismatch and catastrophic forgetting .
Approach: a meta-control system distils on-the-fly abstract representations of states, actions, goals . authors propose a paradigm for autonomous learning driven by pseudo-symbolic abstraction .
Outcome: a meta-control system distils on-the-fly abstract representations of states, actions, goals . a novel approach resolves the domain mismatch problem and lays the groundwork for truly autonomous AI models .
Vision-and-Language Navigation with Analogical Textual Descriptions in LLMs (2025.emnlp-main)

Copied to clipboard

Challenge: Existing zero-shot LLM-based Vision-and-Language Navigation agents either encode images as textual scene descriptions, potentially oversimplifying visual details, or process raw image inputs, which can fail to capture abstract semantics required for high-level reasoning.
Approach: They propose to integrate large language models into embodied AI models by incorporating textual descriptions that facilitate analogical reasoning across images from multiple perspectives.
Outcome: The proposed approach improves the agent’s contextual understanding on the R2R dataset, showing that it can make better decisions based on the LLMs.
How to Mitigate Overfitting in Weak-to-strong Generalization? (2025.acl-long)

Copied to clipboard

Challenge: Experimental results show that weak-to-strong generalization significantly improves PGR compared to naive weak- to-strong . superalignment refers to how humans can align models on tasks beyond human ability to evaluate .
Approach: They propose a framework that elicits the capabilities of strong models through weak supervisors . they propose 'superalignment' to ensure that strong models align with supervisors' intentions .
Outcome: The proposed framework significantly improves quality of supervision signals and quality of input questions compared to naive weak-to-strong generalization .
Distractor Generation in Multiple-Choice Tasks: A Survey of Methods, Datasets, and Evaluation (2024.emnlp-main)

Copied to clipboard

Challenge: Objective questions such as fill-in-the-blank and multiple-choice require examinees to select one valid answer from a set of invalid options.
Approach: They examine distractor generation tasks, datasets, methods, and evaluation metrics for English objective questions.
Outcome: The proposed task is based on fill-in-the-blank and multiple choice questions and is widely utilized in educational settings across various domains and subjects.
Knowledge-Aware Reasoning over Multimodal Semi-structured Tables (2024.findings-emnlp)

Copied to clipboard

Challenge: Existing datasets for tabular question answering focus on text within cells, but real-world data is multimodal, often blending images such as symbols, faces, icons, patterns, and charts with textual content.
Approach: They propose a dataset to assess whether current AI models can perform knowledge-aware reasoning on multimodal structured data.
Outcome: The proposed dataset is a robust benchmark for advancing AI’s comprehension and capabilities in analyzing multimodal structured data.
RaTEScore: A Metric for Radiology Report Generation (2024.emnlp-main)

Copied to clipboard

Challenge: Existing metrics to evaluate the quality of medical reports are limited due to the complexity of clinical free-form texts.
Approach: They propose a new metric to assess the quality of medical reports generated by AI models.
Outcome: The proposed metric is based on a medical NER dataset and trained on NER models . it aligns more closely with human preference than existing metrics, the authors show .
Faithful Persona-based Conversational Dataset Generation with Large Language Models (2024.findings-acl)

Copied to clipboard

Challenge: Existing datasets for training conversational AI models do not sufficiently model their users.
Approach: They propose a generator-critic architecture framework to expand the initial dataset while improving the quality of its conversations.
Outcome: The proposed framework expands the initial dataset while improving the quality of its conversations.
Biases Propagate in Encoder-based Vision-Language Models: A Systematic Analysis From Intrinsic Measures to Zero-shot Retrieval Outcomes (2025.findings-acl)

Copied to clipboard

Challenge: Existing encoder-based vision-language models (VLMs) contain intrinsic biases that manifest in biased outputs.
Approach: They propose a framework to measure intrinsic bias propagation by correlating intrinsic bias with extrinsic bias in zero-shot text-to-image and image-totext retrieval.
Outcome: The proposed framework shows that larger/better-performing models exhibit greater bias propagation, raising concerns given the trend towards increasingly complex AI models.
Fool Me Once? Contrasting Textual and Visual Explanations in a Clinical Decision-Support Setting (2024.emnlp-main)

Copied to clipboard

Challenge: XAI models are being used in safety-critical domains, but their use is limited due to their limited transparency and insufficient model robustness.
Approach: They evaluated visual, natural language and a combination of both modalities to examine how users use them.
Outcome: The proposed model is more robust and transparent than previous models.
Evaluating Reasoning Models for Queries with Presuppositions (2026.findings-acl)

Copied to clipboard

Challenge: Prior work notes that large language models fail to challenge erroneous assumptions and can reinforce users’ misinformed opinions.
Approach: They construct queries with varying degrees of presuppositions spanning health, science, and general knowledge and evaluate several widely-deployed models.
Outcome: The proposed models achieve higher accuracy but fail to challenge a large fraction of false presuppositions.
M-Help: Using Social Media Data to Detect Mental Health Help-Seeking Signals (2025.findings-emnlp)

Copied to clipboard

Challenge: Existing datasets for detecting mental health disorders do not identify individuals actively seeking help.
Approach: This paper introduces a new social media dataset specifically designed to detect help-seeking behavior on social media.
Outcome: The proposed dataset can detect help-seeking behavior on social media . it can address three key tasks: identifying help- seekkers, diagnosing mental health conditions .
Blinded by Context: Unveiling the Halo Effect of MLLM in AI Hiring (2025.findings-acl)

Copied to clipboard

Challenge: Large Language Models (LLMs) and Multimodal Large Language Modells (MLLMs) are increasingly being deployed across a range of domains, including finance, law, peer review, and recruitment.
Approach: They investigated how image-based evaluations are influenced by non-job-related information, including extracurricular activities and social media images.
Outcome: The proposed models exhibit significant halo effects in image-based evaluations while text-based assessments showed more resistance to bias.
PluRule: A Benchmark for Moderating Pluralistic Communities on Social Media (2026.acl-long)

Copied to clipboard

Challenge: Social media are shifting towards community-governed platforms where groups define their own norms.
Approach: They propose a multimodal, multilingual benchmark for detecting 13,371 rule violations across 1,989 Reddit communities . they show that bigger models and increased context provide marginal gains, and universal rules like civility and self-promotion are easier to detect.
Outcome: The proposed model can detect 13,371 rule violations across 1,989 Reddit communities across 2,885 rules in 9 languages.
v-HUB: A Benchmark for Video Humor Understanding from Vision and Sound (2026.acl-long)

Copied to clipboard

Challenge: Humor enriches our daily lives and appears in many forms, from jokes and cartoons to comedies and viral videos.
Approach: They introduce a video humor understanding benchmark to test their ability to understand humor from visual cues.
Outcome: The proposed video humor understanding benchmark is based on a collection of short videos . it features rich annotations and a study of environmental sound that can enhance humor .
Does Theory of Mind Improvement Really Benefit Human-AI Interactions? Empirical Findings from Interactive Evaluations (2026.findings-acl)

Copied to clipboard

Challenge: Existing benchmarks measure ToM capability improvement through story-reading, multiple-choice questions from a third-person perspective, while ignoring the first-person, dynamic nature of human-AI interactions.
Approach: They propose a new paradigm of interactive ToM evaluation with both perspective and metric shifts.
Outcome: The proposed approach improves the performance of four representative LLM enhancement techniques using real-world datasets and a user study.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations